NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Response to ‘Letter to the Editor: on the stability and internal consistency of component-wise sparse mixture regression based clustering’, Zhang et al.

https://doi.org/10.1093/bib/bbac262

Chang, Wennan; Zhang, Chi; Cao, Sha (July 2022, Briefings in Bioinformatics)
Pipeline for characterizing alternative mechanisms (PCAM) based on bi-clustering to study colorectal cancer heterogeneity

https://doi.org/10.1016/j.csbj.2023.03.028

Cao, Sha; Chang, Wennan; Wan, Changlin; Lu, Xiaoyu; Dang, Pengtao; Zhou, Xinyu; Zhu, Haiqi; Chen, Jian; Li, Bo; Zang, Yong; et al (January 2023, Computational and Structural Biotechnology Journal)

Full Text Available
FLUXestimator: a webserver for predicting metabolic flux and variations using transcriptomics data

https://doi.org/10.1093/nar/gkad444

Zhang, Zixuan; Zhu, Haiqi; Dang, Pengtao; Wang, Jia; Chang, Wennan; Wang, Xiao; Alghamdi, Norah; Lu, Alex; Zang, Yong; Wu, Wenzhuo; et al (May 2023, Nucleic Acids Research)

Abstract Quantitative assessment of single cell fluxome is critical for understanding the metabolic heterogeneity in diseases. Unfortunately, laboratory-based single cell fluxomics is currently impractical, and the current computational tools for flux estimation are not designed for single cell-level prediction. Given the well-established link between transcriptomic and metabolomic profiles, leveraging single cell transcriptomics data to predict single cell fluxome is not only feasible but also an urgent task. In this study, we present FLUXestimator, an online platform for predicting metabolic fluxome and variations using single cell or general transcriptomics data of large sample-size. The FLUXestimator webserver implements a recently developed unsupervised approach called single cell flux estimation analysis (scFEA), which uses a new neural network architecture to estimate reaction rates from transcriptomics data. To the best of our knowledge, FLUXestimator is the first web-based tool dedicated to predicting cell-/sample-wise metabolic flux and metabolite variations using transcriptomics data of human, mouse and 15 other common experimental organisms. The FLUXestimator webserver is available at http://scFLUX.org/, and stand-alone tools for local use are available at https://github.com/changwn/scFEA. Our tool provides a new avenue for studying metabolic heterogeneity in diseases and has the potential to facilitate the development of new therapeutic strategies.
more » « less
PLUS: Predicting cancer metastasis potential based on positive and unlabeled learning

https://doi.org/10.1371/journal.pcbi.1009956

Zhou, Junyi; Lu, Xiaoyu; Chang, Wennan; Wan, Changlin; Lu, Xiongbin; Zhang, Chi; Cao, Sha (March 2022, PLOS Computational Biology)
Liu, Jie (Ed.)
Metastatic cancer accounts for over 90% of all cancer deaths, and evaluations of metastasis potential are vital for minimizing the metastasis-associated mortality and achieving optimal clinical decision-making. Computational assessment of metastasis potential based on large-scale transcriptomic cancer data is challenging because metastasis events are not always clinically detectable. The under-diagnosis of metastasis events results in biased classification labels, and classification tools using biased labels may lead to inaccurate estimations of metastasis potential. This issue is further complicated by the unknown metastasis prevalence at the population level, the small number of confirmed metastasis cases, and the high dimensionality of the candidate molecular features. Our proposed algorithm, called P ositive and unlabeled L earning from U nbalanced cases and S parse structures ( PLUS ), is the first to use a positive and unlabeled learning framework to account for the under-detection of metastasis events in building a classifier. PLUS is specifically tailored for studying metastasis that deals with the unbalanced instance allocation as well as unknown metastasis prevalence, which are not considered by other methods. PLUS achieves superior performance on synthetic datasets compared with other state-of-the-art methods. Application of PLUS to The Cancer Genome Atlas Pan-Cancer gene expression data generated metastasis potential predictions that show good agreement with the clinical follow-up data, in addition to predictive genes that have been validated by independent single-cell RNA-sequencing datasets.
more » « less
Full Text Available
Acid-Base Homeostasis and Implications to the Phenotypic Behaviors of Cancer

https://doi.org/10.1101/2022.03.04.482927

Zhou, Yi; Chang, Wennan; Lu, Xiaoyu; Wang, Jin; Zhang, Chi; Xu, Ying (January 2022, Genomics proteomics and bioinformatics)

Acid-base homeostasis is a fundamental property of living cells and its persistent disruption in human cells can lead to a wide range of diseases. We have conducted computational modeling and analysis of transcriptomic data of 4750 human tissue samples of nine cancer types in the TCGA database. Built on our previous study, we have quantitatively estimated the (average) production rate of OH− by cytosolic Fenton reactions, which continuously disrupt the intracellular pH homeostasis. Our predictions indicate that all or a subset of 43 reprogrammed metabolisms (RMs) are induced to produce net protons (H+) at comparable rates of Fenton reactions to keep the intracellular pH stable. We have then discovered that a number of well-known phenotypes of cancers, including increased growth rate, metastasis rate and local immune cell composition, can be naturally explained in terms of the Fenton reaction level and the induced RMs. This study strongly suggests the possibility to have a unified framework for studies of cancer-inducing stressors, adaptive metabolic reprogramming, and cancerous behaviors. In addition, strong evidence is provided to demonstrate that a popular view of that Na+/H+ exchangers, along with lactic acid exporters and carbonic anhydrases are responsible for the intracellular alkalization and extracellular acidification in cancer may not be justified.
more » « less
Full Text Available
Supervised clustering of high-dimensional data using regularized mixture modeling

https://doi.org/10.1093/bib/bbaa291

Chang, Wennan; Wan, Changlin; Zang, Yong; Zhang, Chi; Cao, Sha (July 2021, Briefings in Bioinformatics)

Abstract Identifying relationships between genetic variations and their clinical presentations has been challenged by the heterogeneous causes of a disease. It is imperative to unveil the relationship between the high-dimensional genetic manifestations and the clinical presentations, while taking into account the possible heterogeneity of the study subjects.We proposed a novel supervised clustering algorithm using penalized mixture regression model, called component-wise sparse mixture regression (CSMR), to deal with the challenges in studying the heterogeneous relationships between high-dimensional genetic features and a phenotype. The algorithm was adapted from the classification expectation maximization algorithm, which offers a novel supervised solution to the clustering problem, with substantial improvement on both the computational efficiency and biological interpretability. Experimental evaluation on simulated benchmark datasets demonstrated that the CSMR can accurately identify the subspaces on which subset of features are explanatory to the response variables, and it outperformed the baseline methods. Application of CSMR on a drug sensitivity dataset again demonstrated the superior performance of CSMR over the others, where CSMR is powerful in recapitulating the distinct subgroups hidden in the pool of cell lines with regards to their coping mechanisms to different drugs. CSMR represents a big data analysis tool with the potential to resolve the complexity of translating the clinical representations of the disease to the real causes underpinning it. We believe that it will bring new understanding to the molecular basis of a disease and could be of special relevance in the growing field of personalized medicine.
more » « less
Full Text Available
Spatially and Robustly Hybrid Mixture Regression Model for Inference of Spatial Dependence

https://doi.org/10.1109/ICDM51629.2021.00013

Chang, Wennan; Dang, Pengdao; Wan, Changlin; Lu, Xiaoyu; Fang, Yue; Zhao, Tong; Zang, Yong; Li, Bo; Zhang, Chi; Cao, Sha (December 2021, 2021 IEEE International Conference on Data Mining (ICDM))

In this paper, we propose a Spatial Robust Mixture Regression model to investigate the relationship between a response variable and a set of explanatory variables over the spatial domain, assuming that the relationships may exhibit complex spatially dynamic patterns that cannot be captured by constant regression coefficients. Our method integrates the robust finite mixture Gaussian regression model with spatial constraints, to simultaneously handle the spatial non-stationarity, local homogeneity, and outlier contaminations. Compared with existing spatial regression models, our proposed model assumes the existence a few distinct regression models that are estimated based on observations that exhibit similar response-predictor relationships. As such, the proposed model not only accounts for non-stationarity in the spatial trend, but also clusters observations into a few distinct and homogenous groups. This provides an advantage on interpretation with a few stationary sub-processes identified that capture the predominant relationships between response and predictor variables. Moreover, the proposed method incorporates robust procedures to handle contaminations from both regression outliers and spatial outliers. By doing so, we robustly segment the spatial domain into distinct local regions with similar regression coefficients, and sporadic locations that are purely outliers. Rigorous statistical hypothesis testing procedure has been designed to test the significance of such segmentation. Experimental results on many synthetic and real-world datasets demonstrate the robustness, accuracy, and effectiveness of our proposed method, compared with other robust finite mixture regression, spatial regression and spatial segmentation methods.
more » « less
Full Text Available
Denoising Individual Bias for Fairer Binary Submatrix Detection

https://doi.org/10.1145/3340531.3412156

Wan, Changlin; Chang, Wennan; Zhao, Tong; Cao, Sha; Zhang, Chi (October 2020, Proceedings of the 29th ACM International Conference on Information & Knowledge Management)

Full Text Available
Geometric All-Way Boolean Tensor Decomposition

Wan, Changlin; Chang, Wennan; Zhao, Tong; Cao, Sha; Zhang, Chi (October 2020, Advances in neural information processing systems)

Full Text Available
A data denoising approach to optimize functional clustering of single cell RNA-sequencing data

https://doi.org/10.1109/BIBM49941.2020.9313483

Wan, Changlin; Jia, Dongya; Zhao, Yue; Chang, Wennan; Cao, Sha; Wang, Xiao; Zhang, Chi (December 2020, Proceedings - 2020 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2020)
null (Ed.)
Single cell RNA-sequencing (scRNA-seq) technology enables comprehensive transcriptomic profiling of thousands of cells with distinct phenotypic and physiological states in a complex tissue. Substantial efforts have been made to characterize single cells of distinct identities from scRNA-seq data, including various cell clustering techniques. While existing approaches can handle single cells in terms of different cell (sub)types at a high resolution, identification of the functional variability within the same cell type remains unsolved. In addition, there is a lack of robust method to handle the inter-subject variation that often brings severe confounding effects for the functional clustering of single cells. In this study, we developed a novel data denoising and cell clustering approach, namely CIBS, to provide biologically explainable functional classification for scRNA-seq data. CIBS is based on a systems biology model of transcriptional regulation that assumes a multi-modality distribution of the cells’ activation status, and it utilizes a Boolean matrix factorization approach on the discretized expression status to robustly derive functional modules. CIBS is empowered by a novel fast Boolean Matrix Factorization method, namely PFAST, to increase the computational feasibility on large scale scRNA-seq data. Application of CIBS on two scRNA-seq datasets collected from cancer tumor micro-environment successfully identified subgroups of cancer cells with distinct expression patterns of epithelial-mesenchymal transition and extracellular matrix marker genes, which was not revealed by the existing cell clustering analysis tools. The identified cell groups were significantly associated with the clinically confirmed lymph-node invasion and metastasis events across different patients. Index Terms—Cell clustering analysis, Data denoising, Boolean matrix factorization, Cancer microenvirionment, Metastasis.
more » « less
Full Text Available

« Prev Next »

Search for: All records